Papers with Medical Large Vision-Language Models

2 papers
Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning (2025.acl-long)

Copied to clipboard

Challenge: Existing methods rely on inference-time interventions, which are limited in attention adaptation or require additional supervision.
Approach: They propose a framework for automatic attention alignment tuning that leverages weak labels from SAM and selectively modifies visually-critical attention heads to improve alignment while minimizing interference.
Outcome: The proposed framework outperforms state-of-the-art models on medical VQA and report generation benchmarks.
Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback (2025.acl-long)

Copied to clipboard

Challenge: Existing Medical Large Vision-Language Models (Med-LVLMs) lack visual localization in medical images, which is crucial for abnormality detection and interpretation.
Approach: They propose a medical abnormalities unveiling method based on a Medical Abnormalities Unveiler dataset and propose 'abnormal-aware instruction tuning' and 'abbnormal-Aware Reward' method generates diagnoses based upon identified abnormal areas in medical images.
Outcome: The proposed method outperforms existing medical large vision-language models in identifying and understanding medical abnormalities and improves generalization capability.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations